Back

Trends in Hearing

SAGE Publications

Preprints posted in the last 90 days, ranked by how well they match Trends in Hearing's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Comparison of a Novel Real-World Speech in Noise Auditory Attention Task to Standard Clinical Audiological Metrics

Wade, N. E.; Bormann, B. M.; Mankel, K. M.; Comstock, D. C.; Das, S.; Whittle, R. S.; Brodie, H.; Sagiv, D.; Miller, L. M.

2026-05-29 neuroscience 10.64898/2026.05.28.727658 medRxiv
Top 0.1%
18.8%
Show abstract

Pure tone audiometry (PTA) remains the clinical standard for evaluating hearing ability, yet individuals with similar audiometric profiles often exhibit substantial variability in their capacity to understand speech in everyday listening environments. Growing evidence suggests this variance is related to contributions from cognitive ability and auditory processing that standard threshold measures do not capture. To investigate how PTA, cognitive factors, and demographics such as age jointly predict real-world speech perception, 116 veteran adults 20-70 years old spanning a range of normal to moderate sensorineural hearing losses completed a spatial auditory attention task. Target color words were embedded within naturalistic short-story narratives presented under two conditions: a mono-talker speech-in-quiet (SIQ) condition and a dual-talker speech-in-noise (SIN) condition with a spatially separated competing narrative. Behavioral performance was quantified via color word hit accuracy, reaction time, and comprehension question accuracy. Participants also completed pure tone audiometry, the Montreal Cognitive Assessment (MoCA), and the Speech, Spatial and Qualities of Hearing Scale (SSQ12). Mixed-effects regression models were used to evaluate the contributions of PTA, age, cognitive ability, and self-reported hearing difficulty (SSQ12) to task performance across conditions. Results demonstrate a complex interplay between age, PTA, MoCA, and/or listening condition (SIQ vs. SIN) in predicting identification accuracy, reaction time, and comprehension. Age and condition significantly predicted hit accuracy and reaction time, with older participants showing improved accuracy in quiet but declining accuracy and slower responses in noise. PTA did not emerge as a significant main effect predictor but interacted with cognitive ability and condition to modulate performance, in some cases exhibiting a paradoxical inverse relationship with accuracy dependent on MoCA score. MoCA scores significantly predicted comprehension across conditions, and SIN hit accuracy was positively correlated with SSQ12 scores, validating the task against participants real-world listening experiences. These findings highlight the importance of incorporating cognitive screening and ecologically valid speech perception tasks into audiological assessment to better identify individuals at risk for functional hearing impairment in complex listening environments.

2
Does the method matter? Evaluating the effectiveness, efficiency and ease of hearing-aid gain self-adjustment

Benecke, J.; Whitmer, W. M.

2026-06-12 otolaryngology 10.64898/2026.06.11.26355463 medRxiv
Top 0.1%
18.5%
Show abstract

In conventional hearing-aid personalisation, clinicians cannot hear what their patients hear, and patients cannot often reliably detect or describe what they hear. Self-adjustment avoids this issue but requires user controls that adjust hearing-aid signal processing parameters to be effective, efficient and easy. In this study, we explored (a) the roles of interface complexity and stimulus type in the self-adjustment of hearing-aid gain, and (b) how well individuals can adjust one sound to match another to assess the same interfaces and stimuli. Adult hearing-aid users with mild to moderate symmetrical sensorineural hearing loss repeatedly adjusted the gain (a) to their preference from individual prescription (n = 41) and (b) to match their previous preferences from a random starting point (n = 32) using three interfaces representing different bass/mid/treble configurations and three stimuli (music, speech and speech-in-noise). The large interindividual variability in self-adjusted gains clustered into three patterns of deviation from initial prescription: increased relative bass, overall gain reduction, and close to initial prescription. There were no substantial effects of interface nor stimulus on self-adjustment reliability (median {sigma} = 2.8 dB), whereas absolute sound-matching error increased with increasing interface complexity and centre frequency. Neither individual matching accuracy nor questionnaire responses predicted either self-adjusted gains or reliability. Overall, these results show that many - but not all - hearing-aid users can adjust gains with reasonable reliability, and while it can be difficult to predict the behaviour from the individual, the individual applies a similar self-adjustment behaviour across different interfaces and stimuli.

3
When feedback backfires: investigating neurofeedback effects in a closed-loop auditory attention decoding paradigm

Rotaru, I.; Geirnaert, S.; Heintz, N.; Bertrand, A.; Francart, T.

2026-04-30 neuroscience 10.64898/2026.04.28.721343 medRxiv
Top 0.1%
15.3%
Show abstract

Selective auditory attention decoding (AAD) enables tracking which of multiple concurrent speakers a listener attends to and is a key building block for neuro-steered hearing devices. While AAD integrated in a closed-loop system with real-time neurofeedback (NFB) is hypothesized to improve decoding through neural adaptation and error-correction behaviour, the short-term behavioral and algorithmic impact of such a bilateral human-machine interaction remains poorly understood. Here we evaluated the effects of NFB on AAD accuracy and user experience in a single-session AAD paradigm with online NFB involving nineteen participants. They performed a selective listening task with enforced attention switches across four conditions: open-loop (OL), closed-loop with auditory gain feedback (CLA), closed-loop with visual feedback (CLV), and a condition with pseudo-auditory gain control (psCLA) decoupled from the participants individual neural activity. AAD was performed online using both subject-specific and subject-independent linear decoders on 5 s sliding windows, followed by Hidden Markov Model post-processing. Online analysis showed comparable decoding performance across all conditions. However, offline posthoc analysis using subject-independent decoders revealed that AAD accuracy in the CLA condition was significantly lower than in the OL baseline. Subjectively, participants reported that CLA was significantly more distracting and required higher switching effort. Crucially, a causal analysis of the psCLA condition found no robust evidence that higher audio gains inherently improve decoding accuracy. Our results demonstrate that within a single-session paradigm with rapidly varying feedback cues, auditory neurofeedback may degrade AAD performance by increasing cognitive load and distraction. These findings suggest that suboptimal feedback can impede rather than facilitate learning. We conclude that more accurate and stable decoders and longitudinal, multi-session training protocols are likely essential prerequisites for achieving beneficial neurofeedback effects in closed-loop auditory attention systems.

4
Development and clinical application of a consonant confusion task to evaluate hearing aid benefit

Hajicek, J.; Harris, S. E.; Neely, S. T.

2026-04-24 otolaryngology 10.64898/2026.04.23.26351598 medRxiv
Top 0.1%
12.6%
Show abstract

PurposeThis research sought to develop a low-cognitive-load speech-in-noise test based on consonant confusions with the potential for assessing hearing-aid benefit. MethodsVowel-consonant-vowel (VCV) stimuli with added speech-shaped noise were presented as a closed-set consonant identification task. Initially, consonant-confusion matrices were used to select, from a larger set of consonants and vowel contexts, a set of ten consonants and associated signal-to-noise ratios (SNR) that were sensitive to hearing loss. The sensitivity of the qVCV test to hearing loss was validated by comparing predicted pure-tone average (PTA) hearing thresholds with their audiometric PTA. Clinical viability of the qVCV test was assessed by comparisons to the QuickSIN test. Hearing-aid benefit was assessed by comparing test scores in unaided and aided conditions. ResultsThe consonants most sensitive to hearing loss were /b d g t k v z s [esh] n/ in the vowel context /[a]/. A cross-validated prediction of PTA had a mean-absolute error of 5.7 dB. The repeatability of qVCV at 50 trials was equivalent to the QuickSIN average of two lists. Hearing-aid benefit was quantified as a decibel reduction in hearing loss. ConclusionsqVCV and QuickSIN performed similarly when test times are equated. The advantages of qVCV include lower cognitive demand, fewer learning effects, and automated scoring. PTA predicted by qVCV which greatly exceeds audiometric PTA may indicate either cognitive deficits or cochlear neural degeneration. The qVCV quantification of hearing-aid benefit may have clinical value.

5
Auditory Working Memory Mediates the Relationship between Musical Sophistication and Speech-in-noise Perception

Colak, H.; Benzaquen, E.; Guo, X.; Lad, M.; Sedley, W.; Griffiths, T. D.

2026-05-13 neuroscience 10.64898/2026.05.13.724783 medRxiv
Top 0.1%
9.8%
Show abstract

Understanding speech in noisy environments (SPIN) is an important everyday ability, and engaging in musical activities has been proposed as a factor that may support this ability. However, the cognitive mechanisms underlying a potential musical advantage in SPIN perception remain unclear. Here we investigated whether musical sophistication is associated with better SPIN perception in a large population-based sample, and whether this relationship is mediated by auditory working memory (AWM), verbal working memory (VWM), or non-verbal intelligence. We recruited 203 participants and measured SPIN perception at both word and sentence levels. Musical sophistication was assessed using the Goldsmiths Musical Sophistication Index (Gold-MSI). AWM was measured using delayed matching of tone frequency or the modulation rate of amplitude modulated white noise, VWM was based on backward digit span task, and non-verbal intelligence used matrix reasoning. Mediation analyses revealed that AWM fully mediated the relationship between musical sophistication and SPIN perception, whereas VWM showed no mediation effect. Non-verbal intelligence showed a partial mediating effect. Additional control analyses using structural equation modelling revealed that the indirect effect through AWM remained significant after accounting for age, hearing thresholds, and non-verbal intelligence. Together, these findings suggest that individuals with greater musical sophistication demonstrate better daily life listening abilities, and that superior auditory working memory may be the key cognitive mechanism underlying this advantage.

6
Auditory Working Memory and Sound Segregation Ability Predict Speech-in-Noise in Adult Cochlear Implant Users

Colak, H.; Guo, X.; Benzaquen, E.; Gurusiddappa, M.; Banerjee, A.; Choi, I.; Sedley, W.; Griffiths, T. D.

2026-06-09 neuroscience 10.64898/2026.06.05.730315 medRxiv
Top 0.1%
8.2%
Show abstract

ObjectivesOutcomes following cochlear implantation vary substantially across adult recipients, and the cognitive and perceptual factors contributing to this variability are not fully understood. This poses a challenge for developing strategies to improve cochlear implant outcomes, as such approaches require a clearer understanding of the mechanisms underlying individual listening difficulties. In this study, we investigated auditory cognitive measures in cochlear implant (CI) users to further elucidate the origins of this variability. DesignThirty-seven adult cochlear implant users completed measures of auditory cognition, comprising auditory working memory (AWM) and sound segregation ability, measured using an auditory figure-ground task (AFG), as well as measures of peripheral temporal and spectral processing, comprising the temporal modulation detection threshold (TMDT) and spectral ripple discrimination threshold (SRDT). Speech perception outcomes were assessed using word-in-noise (WIN) and sentence-in-noise (SIN) tasks. Separate multiple linear regression models evaluated the unique contribution of the auditory cognition measures to WIN and SIN performance, after accounting for the peripheral measures. ResultsBoth regression models explained a substantial proportion of variance in speech-in-noise outcomes (WIN: adjusted R{superscript 2} = 0.55; SIN: adjusted R{superscript 2}=0.57, both p < 0.001). For WIN performance, AFG and AWM were significant predictors. A similar pattern was found for SIN performance, where lower AWM ability and poorer AFG segregation were linked to poorer sentence listening in noise. No significant effects of spectral ripple discrimination or temporal modulation detection were observed in either model, even though both were significantly correlated with WIN performance. ConclusionsThese findings indicate that auditory working memory and sound segregation ability are robust predictors of speech-in-noise outcomes in adult cochlear implant users, across both word- and sentence-level measures. Together, the results may help explain why speech-in-noise outcomes remain highly variable among CI users, even when basic sensory encoding abilities are taken into account. Incorporating measures of auditory working memory and fundamental sound segregation may therefore improve outcome prediction and help in developing more individualised rehabilitation strategies.

7
Age-related changes in acoustic cue use for speech-in-speech perception

Fish, E.; DiNino, M.

2026-06-22 otolaryngology 10.64898/2026.06.17.26355866 medRxiv
Top 0.1%
7.9%
Show abstract

Acoustic cues such as pitch and spatial location allow listeners to attend to a target speaker and ignore competing talkers, aiding speech recognition in background noise. Diminished ability to utilize acoustic cues for speech stream segregation may thus contribute to older adults' challenges hearing in noise. Adults aged 18-74 completed a speech-in-speech identification task with three conditions containing 1) only pitch cues (fundamental frequency), 2) only spatial cues (interaural time differences; ITDs), and 3) both pitch and spatial cues for segregating a target talker from competing talkers. Hearing thresholds at standard and extended high frequencies (EHFs), auditory brainstem responses (ABRs), and digit span scores were acquired to examine the influence of sensory and cognitive factors on use of each acoustic cue for speech-in-speech recognition. Significant differences were observed between cue condition scores indicating that use of the available cue(s) drove performance. ABR metrics were not a significant predictor but digit span scores significantly predicted scores on all three cue conditions. Working memory abilities therefore set a baseline for participants' speech-in-speech recognition regardless of the acoustic content. Hearing thresholds at standard frequencies significantly predicted scores on the Pitch condition. EHF hearing thresholds better predicted Spatial and Both Cue condition performance, suggesting that EHF thresholds represent auditory processing important for coding ITDs. Age group analysis revealed that older adults (aged 40+) performed significantly more poorly on all cue conditions of the speech-in-speech recognition task relative to younger adults. Age-related changes in auditory sensory processing may therefore impair older adults' speech-in-noise perception by reducing their ability to use acoustic cues for segregating target and competing speech.

8
Context effects in pitch discrimination reflect response bias not sensory bias

Dirks, C. E.; Guest, D. R.; Oxenham, A.

2026-07-03 neuroscience 10.64898/2026.07.02.735981 medRxiv
Top 0.1%
7.7%
Show abstract

Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.

9
Automated auditory brainstem response peak estimation using a convolutional neural net

Marrone, J. P.; Ziliak, M. C.; Bartlett, E. L.

2026-07-06 neuroscience 10.64898/2026.06.30.735643 medRxiv
Top 0.1%
7.4%
Show abstract

Auditory brainstem responses (ABRs) are a core part of objective functional evaluations of hearing sensitivity and subcortical auditory transmission. Manual assessments of ABR waveforms are still a primary means by which thresholds and peak amplitudes and latencies are measured, which is time-consuming and prone to user variability. Automated methods have offered promising alternatives for ABR classification, but they have sometimes been limited in accuracy or robustness. Here, we developed and tested a supervised convolutional neural network (CNN) based ABR peak classifier that works across sound levels and sound frequencies that can be run quickly on a personal computer using single or dual-channel ABR inputs. For ABR peaks I, III, IV, and V, the classifier achieved over 95% accuracy. High accuracy was maintained even after noise-exposure causing temporary or permanent threshold shifts, and over 90% of peaks were within 0.041 ms (1 sample) of the manually identified peak. Only a few hundred samples were needed to train the network, making it widely amenable to smaller data studies or where the number of subjects or sessions may be low.

10
Neural Tracking of Speech Envelope as an Index of Spatial Release from Masking

Galeano-Otalvaro, J.-D.; Dieudonne, B.; Francart, T.; Wouters, J.

2026-07-02 neuroscience 10.64898/2026.06.29.734758 medRxiv
Top 0.1%
6.2%
Show abstract

Understanding speech in noisy environments relies strongly on binaural cues such as interaural time differences (ITDs) and interaural level differences (ILDs), which support spatial hearing and the segregation of competing sound sources. When these cues are degraded, listeners experience substantial difficulty in complex acoustic environments. Behavioural measures of binaural benefit, such as binaural masking level differences (BMLDs), binaural intelligibility level differences (BILDs), and spatial release from masking (SRM), are well established in normal-hearing (NH) listeners, but they require an active behavioural response. Neural speech tracking using electroencephalography (EEG) has emerged as a promising approach for quantifying neural processing of continuous speech, yet its sensitivity to spatial hearing cues remains insufficiently characterised. In this study, we investigated the neural correlates of spatial release from masking in NH listeners using EEG-based neural speech tracking. Nineteen participants listened to continuous Dutch speech stories presented with masking noise under two spatial configurations, collocated (S0N0) and spatially separated (S0N90), across multiple signal-to-noise ratios (SNRs). Neural tracking of the speech envelope was quantified using both envelope reconstruction and temporal response function (TRF) analyses. Spatial separation enhanced neural tracking of the target speech envelope, particularly at challenging SNRs where behavioural SRM was also observed. TRF analysis further revealed condition-dependent morphologies, including increased amplitudes and decreased latencies of late cortical components consistent with spatial unmasking effects. These neural differences were most pronounced at low SNRs, where spatial cues provide the greatest perceptual benefit. Together, these findings demonstrate that neural speech tracking captures cortical signatures of spatial unmasking and closely reflects behavioural improvements in speech understanding. Establishing these relationships in NH listeners supports the development of objective neural measures for evaluating binaural benefit in difficult-to-test populations.

11
Auditory Profiles in Tinnitus are Age-Dependent: Electrophysiological and Behavioral Evidence

Devolder, P.; Keppler, H.; Dhooge, I.; Verhulst, S.

2026-06-03 neuroscience 10.64898/2026.06.02.729537 medRxiv
Top 0.1%
5.6%
Show abstract

Tinnitus is commonly associated with hearing loss, yet it can also occur in individuals with clinically normal audiometric thresholds. This dissociation has led to the hypothesis that hidden sensorineural hearing loss underlies tinnitus in audiometrically normal-hearing individuals. However, identifying such subclinical deficits non-invasively is challenging because audiometric measures are influenced by age-related changes and interactions among sensorineural processes. In this study, we disentangled the contributions of tinnitus, age, and hearing status to sensorineural encoding and speech perception. We included 113 participants, divided into age- and hearing-status-matched groups with and without tinnitus, and assessed them using otoacoustic emissions, auditory evoked potentials, auditory reflex measurements, and behavioral tasks of speech perception. This design enabled a rigorous evaluation of whether hidden sensorineural deficits underlie tinnitus. Age and hearing status had substantial effects on objective measures of sensorineural function, whereas tinnitus-related effects were subtle and age specific. Young adults with tinnitus and normal audiometric thresholds exhibited enhanced auditory brainstem responses, elevated envelope following responses, and better vowel discrimination. In contrast, middle-aged adults with tinnitus showed no such enhancements and demonstrated poorer speech-in-noise performance. Correlation analyses revealed a tinnitus-related shift toward greater reliance on central auditory processing, compared with the predominantly peripheral associations observed in controls. The middle ear muscle reflex was unaffected by tinnitus but was correlated with hyperacusis-related parameters. Together, these findings suggest distinct tinnitus-related auditory profiles across the lifespan: neural enhancement and improved vowel discrimination in young adults, versus degraded sensorineural encoding and reduced speech intelligibility in middle-aged adults. Significance StatementTinnitus affects a significant portion of the population, yet its underlying origins are still unclear. While hearing loss is a common cause, individuals with tinnitus may also have normal hearing thresholds. This suggests that subtle sensorineural damage may also play a role. This study critically investigates tinnitus-, age-, and hearing-related sensorineural encoding using non-invasive electrophysiological measures, auditory reflexes, and speech perception tasks in carefully matched participant groups. The study reveals distinct tinnitus-related auditory profiles throughout the lifespan; including enhanced sensorineural processing in young adults and degraded encoding with impaired speech perception in middle-aged adults. These findings provide critical insight into the mechanisms underlying tinnitus and offer objective markers for future research on tinnitus diagnosis and treatment

12
Effects of Aging, Hearing Loss, and Co-Activation on the Middle Ear Muscle Reflex and Medial Olivocochlear Reflex

Devolder, P.; Deloche, F.; Thienpont, M.; Keppler, H.; Verhulst, S.

2026-04-28 otolaryngology 10.64898/2026.04.27.26351829 medRxiv
Top 0.1%
5.6%
Show abstract

The middle ear muscle reflex (MEMR) and medial olivocochlear reflex (MOCR) are increasingly studied for their role in suprathreshold auditory processing. However, recording these reflexes in humans is potentially complicated by age-related (sub)clinical hearing loss and co-activation. This study investigates (1) the influence of age-related (sub)clinical hearing loss, (2) methodological differences between conventional and wideband MEMR techniques, and (3) how MEMR activation contaminates MOCR recordings. Three test groups were included: young normal-hearing adults, middle-aged normal-hearing adults, and middle-aged adults with audiometric hearing loss. Cochlear status and neural encoding was assessed using distortion-product otoacoustic emissions (DPOAEs) and envelope following responses (EFRs). MEMR recordings were compared using conventional tonal stimuli and wideband stimuli. MOCR was recorded at elicitor levels of 60 and 75 dB to evaluate MEMR co-activation. MEMR was related to age, suggesting sensitivity to subclinical cochlear damage. Wideband stimuli were beneficial as elicitor (noise vs. tone), while changing the probe stimuli added no significant benefit (click vs. tone). MOCR strength did not correlate with age-related subclinical hearing, suggesting that MOCR measurements may reflect efferent function relatively independently of afferent sensorineural status in audiometric normal hearing subjects. However, reliable recordings were challenging in participants with audiometric hearing loss due to poor OAE baselines. MEMR co-activation was detectable in the click response and could alter MOCR-induced suppression. These findings suggest that, in cases of normal hearing thresholds, MEMR amplitude may be a marker of subclinical cochlear damage and MOCR measurements may more specifically reflect efferent function. Clinical measurements can be improved using broadband stimuli, accounting for outer-hair-cell damage, and defining criteria for reflex co-activation.

13
Stability of phoneme-related potentials across testing sessions and stimulus presentation conditions

Guo, Z.-c.; McFarlane, K.; McHaney, J. R.; Choksi, I.; Feeney, M.; Preston, L.; Chandrasekaran, B.

2026-05-29 neuroscience 10.64898/2026.05.26.727924 medRxiv
Top 0.1%
5.5%
Show abstract

ObjectivesObjective and ecologically valid measures of speech processing can complement conventional audiologic assessments. Phoneme-related potentials (PRPs), derived by averaging listeners electroencephalography (EEG) responses time-locked to phonemes in continuous speech, have emerged as a promising approach for capturing cortical processing of speech in naturalistic listening conditions. Importantly, PRPs reveal speech perception challenges even when conventional audiograms are clinically normal, positioning them as a promising neural marker for suprathreshold listening difficulties that standard audiometry often misses. As a critical step toward clinical translation, this study examined the extent to which PRP-derived measures remain stable across real-world contexts relevant to clinical implementation, including monaural versus binaural presentation, stimulus intensity level, and repeated testing sessions. The study also assessed cortical tracking of lower-level speech acoustics to determine whether the PRP findings could be attributed to acoustic processing. DesignEEG was recorded from 18 young adults with normal hearing as they listened to audiobook speech presented monaurally or binaurally at 60 or 75 dB across two sessions separated by approximately one week. Neural differentiation of phoneme manner-of-articulation classes (vowels, nasals/approximants, fricatives, and stops) in PRPs was quantified using two measures: an F-statistic reflecting between-manner relative to within-manner variability, and classification accuracy from a machine-learning model trained to predict manner class from PRPs. Temporal response function modeling assessed neural tracking of continuous acoustic envelope and onset features of the audiobook speech. ResultsNeither PRP-derived measure of manner differentiation showed significant effects of session, presentation modality, intensity level, or their interactions. Intraclass correlation analyses further indicated moderate-to-good reliability across all three factors. In contrast, neural tracking of the acoustic envelope and acoustic onsets was stronger under binaural than monaural presentation, with binaural presentation eliciting more pronounced cortical responses to the envelope. ConclusionsPRP-derived measures remained relatively stable across modest procedural variations that are common in clinical testing contexts, positioning PRPs as a potent objective index of naturalistic speech processing. This stability may reflect cortical processing of abstract, linguistically relevant speech categories and suggest that PRPs provide complementary information beyond audiologic assessments of peripheral auditory functions and EEG measures that primarily capture lower-level acoustic processing.

14
Physiological Markers of Auditory Situational Awareness in Complex Spatialised Scenes

Sztandera, J.; Poole, K. C.; Shiell, M. M.; Picinali, L.; Chait, M.

2026-05-29 neuroscience 10.64898/2026.05.27.728177 medRxiv
Top 0.1%
5.5%
Show abstract

Detecting changes in acoustic environments is essential for situational awareness. It remains unclear whether the spatial location of such changes modulates automatic orienting and arousal mechanisms. We measured pupil dilation, pupil dilation rate, and microsaccade rate while listeners (n=25) heard complex, spatialized auditory scenes rendered over headphones using individualized HRTFs. Participants were naive to the critical manipulation: the appearance of a new source from one of five locations: front, left, right, back, or above. A subsequent localization task assessed perceptual spatial uncertainty. Behaviorally-irrelevant source appearances elicited a cascade of ocular responses. Microsaccadic inhibition emerged from [~]85ms after change onset, and was broadly comparable across locations, suggesting a location-invariant early orienting response to auditory change. Pupil dilation rate increased from [~]200ms, followed by a phasic pupil dilation response from [~]400ms, indicating engagement of arousal-related systems. Pupil responses were modulated by source location: changes from front/left/right elicited larger dilation than changes from above, with back responses showing a similar but weaker reduction. Behavioral localization revealed substantial confusion for front/back/above locations. However, this did not mirror the physiological data, as front sources elicited pupil responses comparable to lateral sources. These findings demonstrate that complex auditory scene changes recruit oculomotor and autonomic systems even outside the focus of task relevance. They further suggest a dissociation between early, location-invariant attentional capture indexed by microsaccadic inhibition and later, location-sensitive arousal indexed by pupil dilation. Spatial biases in auditory situational awareness therefore appear to emerge after initial change detection, shaping arousal and behavioral performance rather than the earliest orienting response.

15
Hearing Lips and Seeing Voices After Fifty Years: A Large-Scale McGurk Illusion Dataset for Audiovisual Speech Research

Wang, Z.; Li, G.; Yu, Y.; Wu, J.; Yu, Z.; Meng, Y.; Wang, S.; Dong, C.

2026-06-10 neuroscience 10.64898/2026.06.09.731046 medRxiv
Top 0.1%
4.9%
Show abstract

Efficient face-to-face communication relies on the integration of auditory speech and visual articulatory signals. Over the past five decades, the McGurk illusion has been widely used as an index of audiovisual speech integration. However, substantial variabilities in susceptibility to the illusion across participants and speakers limit its reliability as a stable measure of audiovisual integration ability. Here, we introduce the McGurk illusion dataset (MID), which, to our knowledge, is the largest publicly available McGurk stimulus dataset to date. The MID comprises auditory (N = 400), visual (N = 400), and audiovisual (N = 640) speech stimuli generated from 80 Mandarin speakers and validated through behavioral judgments across 360,900 trials. Using this dataset, we characterized the acoustic and facial articulatory properties of McGurk stimuli, replicated substantial inter-participant and inter-speaker variabilities in illusion susceptibility, and revealed the associations between variations in McGurk illusion rate and the variations in unisensory perception, audiovisual correspondences, and speakers characteristics. Furthermore, the stimulus set enabled systematic comparisons of the reliability of different McGurk illusion-based indices of audiovisual speech integration. Overall, the MID not only provides a standardized resource for investigating audiovisual speech integration and its alterations across populations, but also supports research on speaker normalization, lip-reading, and speech perception.

16
Modulation statistics allow robust prediction of speech recognition accuracy across many words, voices, and natural background sounds.

Clonan, A. C.; Stevenson, I. H.; Escabi, M. A.

2026-04-30 neuroscience 10.64898/2026.04.27.721224 medRxiv
Top 0.1%
4.8%
Show abstract

Although humans excel at speech recognition, recognition accuracy can vary widely due to differences in background environments as well as the speakers voice quality, intonation, and pitch. Predicting when speech recognition will succeed or fail, however, remains an ongoing challenge in hearing research. Here we characterize recognition abilities across a wide range of natural conditions using digits spoken by many male and female talkers of multiple ages with 33 unique backgrounds. Across this diverse set of sounds, speech recognition is most strongly influenced by the spectrum and modulation statistics of the noise. Yet, articulatory features of the speech, including fundamental and formant frequencies, show categorically distinct modulatory effects on accuracy across age, gender, and words. We then show that a low-dimensional model of sound, based on computations in the auditory midbrain, accounts for participants single-trial recognition behavior across voices, words and backgrounds. Thus, speech-in-noise perception across extremely diverse natural conditions depends largely on a simple set of spectrotemporal statistics likely encoded by central neural populations.

17
Sound Localization Is Biased By Simultaneous And Delayed By Preceding Visual Distractors

Rocchi, F.; Haukes, N. C.; van Opstal, A. J.; van Wanrooij, M. M.

2026-05-15 neuroscience 10.64898/2026.05.12.724474 medRxiv
Top 0.1%
4.6%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWVision can shape auditory perception, especially when visual cues occur at different times and locations than sounds. Simultaneous but spatially misaligned lights bias the perceived location of a sound--a phenomenon known as the ventriloquism effect. Temporally misaligned lights can also affect the latency of auditory responses. However, it remains unclear how multiple visual stimuli that differ from sounds in both space and time jointly influence localization behaviour. We investigated how visual distractors, spatially misaligned by 10{degrees}, presented before and/or during a target sound influence localization accuracy and response latency in a rapid head-pointing task. Human listeners localized brief (150 ms) broadband noise bursts with an average root-mean-square error of 5{degrees} and a baseline latency of 252 ms. Simultaneous visual cues induced the ventriloquism effect, in which the perceived sound location was biased by 1.8{degrees}. Response latency also increased by 21 ms (273 ms). Preceding visual stimuli (2 s duration) did not induce a bias, but increased latency by 55 ms (307 ms). Introducing a 200 ms gap between the preceding light and the sound reduced this latency increase to 24 ms (276 ms), still not inducing a significant bias. When we presented both a preceding and a simultaneous light on opposite sides of the sound, localization reflected the bias induced by the simultaneous light (1.8{degrees}) and the latency increase induced by the preceding light (by 48 ms). These findings reveal a dissociation in audiovisual integration: preceding visual stimuli primarily influence when a sound is responded to (latency), while simultaneous stimuli influence where it is perceived (accuracy). This supports causal inference models of multisensory integration and suggests distinct underlying mechanisms for spatial and temporal processing of sounds in sensorimotor circuits.

18
Speech Stream Tracking in 2D: Attention Differentially Enhances Acoustic and Phonemic Encoding Across Spatial Planes

Kenti Kranidioti, V.; Schonwiesner, M.

2026-06-03 neuroscience 10.64898/2026.06.03.729740 medRxiv
Top 0.1%
4.5%
Show abstract

Selective attention enables the flexible allocation of neural resources toward relevant stimuli. In audition, it allows listeners to track a target stream within acoustically complex environments. This raises the question of how strongly auditory attention depends on spatial cues to achieve stream segregation and maintain distinct auditory objects. Evidence shows that attention enhances neural encoding of sound features for streams separated in azimuth. However, it remains unclear whether the same mechanisms apply without binaural cues in elevation, and how attention prioritizes acoustic versus linguistic features under these conditions. To address this, participants listened to arrhythmic streams of digits spoken by the same voice and separated in azimuth or elevation. They attended one stream to detect target numbers while ignoring the other. Neural responses to attended and ignored streams were modelled using envelope and phoneme temporal response functions, allowing comparison of low-level envelope and higher-level phonemic encoding across spatial dimensions. Results revealed distinct feature-weighting profiles across spatial planes. In azimuth, selective attention was primarily supported by enhanced envelope encoding of the target stream at early and middle latencies. In elevation, envelope encoding was reduced, while phoneme encoding exhibited a more widespread attentional modulation. These findings suggest that phonemic representations support selective stream tracking when binaural spatial cues are unavailable, reflecting flexible weighting of acoustic and phonemic information depending on spatial cue availability.

19
Differentiating the Physiological Signatures of Cochlear Synaptopathy and Inner Hair Cell Damage in a Chinchilla Model

Sivaprakasam, A.; Schweinzger, I.; Heinz, M.

2026-05-08 neuroscience 10.64898/2026.05.05.723072 medRxiv
Top 0.1%
4.2%
Show abstract

Aging and noise over-exposure lead to complex mixtures of cochlear degradation that impair the structure and function of outer hair cells, inner hair cells (IHCs), and the cochlear nerve. However, IHC damage and cochlear synaptopathy (CS) remain pathologies "hidden" from the audiogram. This study aimed to identify and differentiate the physiological signatures of these two distinct pathologies using promising non-invasive assays: Envelope Following Responses (EFRs), Auditory Brainstem Response (ABRs), Wideband middle-ear reflexes (WB-MEMRs), and Distortion Product Otoacoustic Emissions (DPOAEs). We utilized chinchilla models of carboplatin-induced (CA) IHC damage (N = 4) and temporary threshold shift (TTS) noise-induced CS (N = 4) to compare the physiological signatures of each pathology. While both groups showed unchanged ABR thresholds two weeks after exposure, EFRs, ABR Wave V/I ratios, and MEMRs showed distinct effects of exposure. Despite non-elevated ABR-derived audiometric thresholds after exposure, both CA and TTS exposure resulted in severe in EFR "peakiness", particularly for sharp, short-duty-cycle stimuli and significant elevations in ABR Wave V/I ratios. However, these findings were less-pronounced in the TTS-exposed animals. WB-MEMR amplitudes were decreased with elevated thresholds in both groups; this effect was more pronounced in the TTS group. Opposite trends in DPOAE amplitudes indicated that while both IHC damage and CS result in similar suprathreshold temporal coding deficits, effects on outer-hair-cell integrity and auditory efferent physiology may differ between the two pathologies. Future work and novel diagnostics should aim to distinguish these specific cochlear pathologies in clinical populations, or at the very least consider their overlap. HighlightsO_LIA multi-metric diagnostic approach was used with chinchilla models of inner-hair-cell (IHC) damage and cochlear synaptopathy (CS). C_LIO_LIIHC damage and synaptopathy both cause suprathreshold deficits "hidden" from the audiogram. C_LIO_LIIHC damage results in more severe temporal envelope coding degradation than does synaptopathy. C_LIO_LIA combination of EFR "peakiness", ABR Wave V/I ratio, and Wideband Middle Ear Muscle Reflex (WB-MEMR) appear to be useful measures for profiling IHC damage and CS. C_LI

20
Tune Out: A randomised controlled trial to investigate the impact of an online program on tinnitus severity, handicap, and psychological symptoms in adults with tinnitus.

Laird, E. C.; Gosbell, D.; Dall'Est, A.; Malicka, A.

2026-07-08 otolaryngology 10.64898/2026.07.05.26357341 medRxiv
Top 0.1%
4.2%
Show abstract

Objective: To evaluate the efficacy, engagement, and usability of Tune Out, an unguided, self-paced online tinnitus management program, for reducing tinnitus severity in adults with tinnitus. Design: A two-arm, parallel-group randomised controlled trial was conducted with Australian adults reporting diagnosed or self-reported tinnitus. Participants were randomised to immediate access to Tune Out or a waitlist control group. Outcomes were assessed at baseline, 6 weeks, and 12 weeks. The primary outcome was tinnitus severity measured using the Tinnitus Functional Index (TFI). Secondary outcomes included tinnitus handicap, psychological symptoms, program engagement, self-efficacy, and usability. Results: Eighty-eight participants were randomised: 43 to the intervention group and 45 to the waitlist control group. The primary outcome analysis included 63 participants at 12 weeks. A significant Group x Time interaction was observed for TFI total score, indicating greater reductions in tinnitus severity over time in the intervention group compared with waitlist control, F(2, 102.57) = 5.95, p = .004, partial 2= .104. Significant effects were also observed for tinnitus handicap, F(2, 106.76) = 4.12, p = .019, partial 2 = .072. Effects on psychological symptoms were less consistent, although anxiety showed a significant Group x Time interaction, F(2, 116.85) = 3.63, p = .030, partial 2 = .059. At 12 weeks, 23.1% of intervention participants achieved a clinically meaningful reduction in tinnitus severity compared with 5.4% of controls. Program use was highly variable, with a median use of 1.10 hours, and 25.6% of intervention participants recording no use. Usability ratings were favourable among respondents, with a mean System Usability Scale score of 73.13. Conclusions: Tune Out demonstrated preliminary efficacy for reducing tinnitus severity and tinnitus handicap compared with waitlist control. Effects on broader psychological symptoms were less consistent. Although usability was rated positively, low and variable engagement highlights the need for strategies to support uptake and sustained use in unguided digital tinnitus interventions.